Papers with visually-rich document understanding

3 papers
Hypergraph based Understanding for Document Semantic Entity Recognition (2024.acl-long)

Copied to clipboard

Challenge: Existing document understanding models focus on entity categories while ignoring the extraction of entity boundaries.
Approach: They propose a hypergraph attention document semantic entity recognition framework which uses hypergraph focus to focus on entity boundaries and entity categories at the same time.
Outcome: The proposed framework can improve the performance of existing models on FUNSD, CORD, XFUND and SROIE.
ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training (2025.coling-main)

Copied to clipboard

Challenge: Recent approaches for visually-rich document understanding use manually annotated semantic groups.
Approach: They propose a new variant of the VrDU task that does not use manually annotated semantic groups.
Outcome: The proposed method improves on the existing methods while sacrificing performance.
ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for visually rich document understanding lack layout-centered knowledge . experimental results show that ERNIE-Layout improves layout awareness .
Approach: They propose a document pre-training solution with layout knowledge enhancement in the whole workflow to learn better representations that combine the features from text, layout, and image.
Outcome: The proposed model outperforms existing models on key downstream tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations